Skip to content

Cluster-aware bootstrap resampling for repeated measures designs - #242

Open
mlotinga wants to merge 8 commits into
ACCLAB:masterfrom
mlotinga:main
Open

mlotinga wants to merge 8 commits into
ACCLAB:masterfrom
mlotinga:main

Conversation

@mlotinga

Copy link
Copy Markdown

Add cluster-aware bootstrapping and permutation tests (cluster_col)

Addresses enhancement feature request issue #241.

Summary

DABEST's bootstrap and permutation test treat every row of id_col (or, for unpaired data, every observation) as an independent sampling unit. Some repeated-measures designs break that assumption: when a participant contributes several paired sets, for example several stimuli rated under each condition. Analysing these designs using DABEST represents pseudoreplication, with the impact that confidence intervals are estimated as too narrow.

This feature adds a cluster_col argument to dabest.load() naming the real independent sampling unit. When it is set, whole clusters are resampled and reshuffled, and intervals are expanded for the number of clusters, so that intervals reflect between-cluster correlation and keep close to nominal coverage. When it is not set, nothing changes.

What changed

  • dabest.load(..., cluster_col=None, cluster_ci_expansion=True): works with unpaired and paired data (id_col identifies the pairs, cluster_col the units they are nested in), shared-control and multi-group idx, proportional (binary) data, delta-delta and mini-meta.
  • Cluster bootstrap: whole clusters are resampled with replacement, stratified by the pattern of groups each cluster appears in, so designs where a cluster spans both groups, designs where clusters are nested within one group, and mixtures of the two are each resampled correctly. Delta-delta resamples clusters jointly across all four groups.
  • Small-sample expansion (on by default): bootstrap intervals are too narrow when there are few independent units, so intervals are expanded for the number of clusters with Hesterberg's (2015) expanded percentile method. Each interval is read further into the tails of the same bootstrap distribution, using t-distribution degrees of freedom based on the number of clusters. For nested or mixed designs, whose groups of clusters are resampled independently, the groups' degrees of freedom are combined with the Welch-Satterthwaite approximation. This applies to percentile and BCa intervals, the baseline error curve, delta-delta and mini-meta. The reported level (ci) is unchanged; the level actually read is reported as ci_expanded, with the unexpanded limits alongside. cluster_ci_expansion=False turns it off.
  • BCa acceleration uses a delete-one-cluster jackknife.
  • Cluster-level permutation test: whole-cluster sign flips for paired data, whole-set swaps for clusters spanning both groups, and cluster reassignment for nested clusters, each with the matching permutation count for the Phipson & Smyth adjustment.
  • Results and API: n_clusters, ci_expanded and unexpanded-limit columns (bca_low_unexpanded, pct_high_unexpanded, bec_bca_low_unexpanded, ...) in .results, which appear only when they apply; ci_expanded in .statistical_tests; TwoGroupsEffectSize and PermutationTest accept control_clusters/test_clusters (and TwoGroupsEffectSize accepts cluster_ci_expansion).
  • Plots:
    • With cluster_col set, each group label reads (N=<observations>, / n=<clusters>) on two lines. In vertical Cumming plots (including two-column Sankey plots), the gap between the raw-data and contrast axes is sized to the labels.
    • Because the plotted bootstrap distribution is unchanged, an expanded interval is drawn in two parts, in estimation plots and forest plots. The unexpanded interval, the nominal interval of the plotted distribution, is the usual thick bar, and the expansion beyond it is a thinner line. The thin line can be styled with the new contrast_expanded_errorbar_kwargs in .plot() and expanded_errorbar_kwargs in forest_plot().
  • Validation and warnings: clear errors for a missing column, missing cluster labels, a cluster column that duplicates x, y or an idx group, and pairs whose rows carry different cluster labels; a warning when too few clusters remain for any bootstrap
    interval to be reliable.

Point estimates are unaffected. The parametric and rank-based tests in .statistical_tests are unchanged and still treat observations as independent.

Accuracy checks

  • With one row per cluster, every cluster-aware resampling path reproduces the existing bootstrap bit for bit, for continuous and proportional data.
  • Cluster-bootstrap standard errors matched closed-form cluster-sampling theory to within a few percent in simulation.
  • Coverage of the shipped code was checked against known true effects for every effect size. The designs covered were paired, balanced and unbalanced nested, mixed and binary, with 8 to 16 clusters and 400 to 500 datasets per cell. With at least 6 clusters, and at least 4 in each independently resampled group, expanded intervals covered 92 to 98% at a nominal 95%. Unexpanded cluster-bootstrap intervals covered 80 to 95%, and naive intervals less.
  • The new tutorial reproduces a coverage check end to end. Across 200 simulated datasets, the naive 95% interval covers the true effect about 82% of the time, the unexpanded cluster-aware interval about 90%, and the default expanded interval about 95%.

Documentation

  • New tutorial, Cluster-Robust Bootstrap for Repeated Measures:
    • a worked example where, for the same data and point estimate, the naive interval lies entirely above zero while the cluster-aware interval spans it;
    • the coverage simulation above;
    • a section on the small-sample expansion and how plots draw it.
  • Shared Control & Repeated Measures tutorial: a cluster_col walkthrough.
  • Plot Aesthetics tutorial and show_baseline_ec docstring: the baseline error curve is always an unpaired self-comparison, so with cluster_col it can be much wider than the real paired contrasts.
  • Docstrings for load(), plot(), forest_plot(), show_sample_size, TwoGroupsEffectSize and PermutationTest, plus a CHANGELOG entry.

Limitations

  • With fewer than about 6 clusters, or fewer than 4 in any independently resampled group, no bootstrap interval is reliable, even after expansion; DABEST warns in that case.
  • In unbalanced or mixed designs whose smallest group holds only 4 to 6 clusters, expanded intervals can still cover 1 to 3 points below nominal. Using the smallest group's degrees of freedom instead would over-cover by similar amounts, with intervals about 10% wider.
  • Mini-meta resamples each experiment independently, so clusters shared across experiments are not resampled jointly. The delta-delta permutation p-value likewise combines per-comparison permutations.

Testing

  • New nbs/tests/test_cluster_bootstrap.py: 34 tests. They cover:
    • the resampling machinery and equivalence with the existing bootstrap;
    • the permutation schemes;
    • the small-sample expansion: maths, strata, defaults, opt-out, edge cases, delta-delta, mini-meta and every effect size;
    • the unexpanded limits and their two-part drawing in estimation and forest plots;
    • validation, plot labels and measured label clearance.
  • Full suite: 468 pytest tests pass, including all image comparisons unchanged.
    Every test, API and tutorial notebook executes under nbdev-test. All notebooks are clean and in sync with the library.

Backward compatibility

cluster_col defaults to None. Every new code path is gated on it being set, including the expansion, the two-part interval, the plot-label and layout changes, and the new result columns.

Bug (fixed): for unpaired data where the same clusters appear in both groups, the permutation test reshuffled labels within each cluster. That tests the strong null of no effect in any participant, so participant-specific effects that average to zero still caused rejections: 20% at the 5% level. It now swaps each cluster's whole control and test sets, the direct analogue of the paired sign flip. After the fix it rejects 3% at the 5% level and 10% at the 10% level. See dabest/_effsize_objects.py:1885.
Baseline error curve documentation did not mention the details that make its use questionable to paired or clustered analyses.

show_baseline_ec docstring (dabest/_effsize_objects.py:1584): now explains what the curve actually computes, states plainly that it's always an unpaired self-comparison regardless of paired, and calls out the scale mismatch with cluster_col.

cluster_col docstring (dabest/_api.py:92): adds a pointer to that same caveat, since a user reading about cluster_col alone wouldn't otherwise know it affects the baseline curve disproportionately.

Plot Aesthetics tutorial (08-plot_aesthetics.ipynb): rewrote the "Baseline error curve" section and added a new subsection with a worked, executed example. It builds a 15-participant, 3-pairs-each dataset and plots the same comparison with id_col alone versus id_col plus cluster_col, side by side. The real paired contrast stays essentially unchanged between the two; the baseline curve at the "Control" position visibly widens, from a 95% width of about 1.2 to about 2.0 in this example. That's the same phenomenon you flagged, reproduced deliberately and small enough to read at a glance, rather than buried in a large real dataset.

Changelog: added a Documentation entry alongside the existing cluster-aware bootstrap entry.
Added new tutorial to demonstrate the use and value of the cluster-aware bootstrap feature, and added hooks in associated notebooks.
Ensure plot labels reflect both observations and cluster sample sizes.
…ter sample sizes

Hesterberg expanded percentile adjustment to expand coverage for small numbers of clusters. Active when cluster_col is assigned, opt-out argument also available.
Degrees of freedom adjustment for percentile expansion switched from Hsu to Welch-Satterthwaite to avoid over-conservatism. Plots also now indicate expanded intervals with thinner lines, also editable.
@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

Extra gap added to accommodate cluster sample size in plots needed to be rationalised.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant